rl: add the elastic resource benchmark contract, gating, evidence and paired work - #66
Merged
Merged
Conversation
…red work Implements the GPU-free half of the rl-elastic-resource-benchmark change: versioned study manifest with canonical hash and calibration/formal modes, config/edge/pool validation gated by runtime capability attestation, evidence index with tamper-rejecting resume, dry-run plan CLI, layered results and report, and paired splits/schedules/scenarios that reuse the legacy benchmark_rl pairing without changing its CLI. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
michaellchung
added a commit
to michaellchung/yeto
that referenced
this pull request
Sep 29, 2026
…ed fingerprint, pause audit fixes) - tasks 1.4/1.5/1.6 unchecked: implemented + CPU tests pass, dependencies (1.2/1.3, GPU X9) not met; progress.md updated. - accept spec spelling "serial-colocated" as alias of "colocated-serial" across manifest, attestation, EngineCapabilities and ExecutionProfile. - fingerprint_rejection fails closed when the study fingerprint is None or "unresolved"; agentenv#66 tests pin a fingerprint where they expect support. - pause-audit.md: read_loop line 1259, decoupled budget mode max_reconnects=None, BUDGET_DONE lease only fatal in the collect_budget_reports window in learner-budget mode, non-budget decoupled marked unproven, 450 s documented as policy not a code limit. - pause_audit.py: budget derived from quorum_timeout_s x margin; unused constants removed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
michaellchung
added a commit
to michaellchung/yeto
that referenced
this pull request
Sep 29, 2026
…alized (Verda not supported), DynaResize hypotheses with page citations runtime_manifest.py collects commits/import paths/versions/fork interfaces in the image and refuses certification on pin mismatch or a missing interface for a declared capability. gpu-plan.md section 8 maps experiments to the existing launcher (Nebius sky, Modal runner), agentenv#66 pool identity, per-rental record and mechanical cleanup; notes the Modal H100 (no '!') launcher gap. DynaResize (arXiv:2607.22614) turned into H1-H6 with page numbers, miles adaptation points and negative-result handling; no paper constants adopted. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The GPU-free half of the
rl-elastic-resource-benchmarkstudy harness:scripts/benchmark_rl_elastic.pyplans and reports studies comparing fixed GPU partitions (trainer / rollout / standby) against scheduled and automatic resizing inside one RL island.What lands
planCLI, layered results and areportthat exits non-zero while the matrix is incompletebenchmark_rlpairing throughyeto/rl/elastic_benchmark/legacy.pyscripts/benchmark_rl.py, its CLI and its result fields are unchanged.No GPU runner ships here. None of the four commands loads a model, imports torch/ray/miles, or creates cloud resources.
Conflicts
None — every file is new (
yeto/rl/elastic_benchmark/, the script, the test module, the doc). Nothing existing is touched.Spec
The planning contract lives in the miles repository under
openspec/changes/rl-elastic-resource-benchmark/, asdocs/RL_ELASTIC_BENCHMARK.mdstates; there is deliberately no OpenSpec change for it in this repo.Tests
19 pass locally (
tests/test_rl_elastic_benchmark.py).🤖 Generated with Claude Code